Back

Microbiology Resource Announcements

American Society for Microbiology

Preprints posted in the last 90 days, ranked by how well they match Microbiology Resource Announcements's content profile, based on 25 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Sequencing, Chromosome-scale Assembly, and Annotation of the Genome of the Halophilic Nanoflagellate Halocafeteria seosinensis

Gallot-Lavallee, L.; Haro, R.; Jerlstrom-Hultqvist, J.; Tymoshenko, D.; Roger, A.; Archibald, J. M.

2026-06-30 genomics 10.64898/2026.06.25.734631 medRxiv
Top 0.1%
18.4%
Show abstract

Compared with bacterial and archaeal extremophiles, single-celled eukaryotes living in extreme habitats are understudied and underrepresented in genomic databases. An exception is the obligately halophilic stramenopile Halocafeteria seosinensis strain EHF34. A transcriptome-focused analysis of this extremophilic protists revealed the importance of organic osmolyte regulation and transport in its adaptation to hypersaline environments. However, genomic resources for H. seosinensis are currently limited to a highly fragmented assembly generated by short-read sequencing, which has hindered further investigation of the genome biology and evolution of this fascinating organism. Here, we used long-read Oxford Nanopore sequencing to generate a highly contiguous, chromosome-scale genome assembly for H. seosinensis. The assembly is 38.8 megabase pairs (Mbp) in size and contains 60 nuclear contigs, making it the most contiguous genome for a member of the order Bicosoecida. Approximately 19% of the genome is comprised of transposable elements. Of the 11,684 predicted protein-coding genes, many appear to be associated with DNA mobility-related functions, and several may be linked to adaptation to a hypersaline environment. Analysis of the H. seosinensis long-read genome assembly presented herein will facilitate our understanding of the ways in which protists have adapted to extreme environments. SignificanceHalocafeteria seosinensis is an extremophilic protist adapted to hypersaline environments. Previous analyses of a transcriptome and short-read draft genome assembly for this organism provided insights into the molecular mechanisms underlying osmotic regulation, which facilitate its adaptation to high-salt conditions. However, the lack of contiguity and quality of the draft assembly prevented the characterization of complex genomic regions, including transposable elements and viral insertions, as well as genomic comparisons with related species. Here we present a highly contiguous, chromosome-scale genome assembly for H. seosinensis that enables accurate gene prediction, detailed analysis of repeat content, and comparative genomic analysis. This long-read genome assembly will serve as a valuable resource for studying one of the few tractable halophilic protists sequenced to date.

2
Reysenbachia aerophila gen. nov., sp. nov., a facultatively anaerobic, hydrogen-oxidizing, thermophilic bacterium isolated from Kuirau Park, Rotorua, New Zealand

Marshall, M. E. A.; Stott, M. B.; Welford, H. E.; Lagutin, K.; Mitchell, K. A.; Carere, C. R.

2026-06-15 microbiology 10.64898/2026.06.14.732183 medRxiv
Top 0.1%
12.9%
Show abstract

A facultatively anaerobic, hydrogen-oxidizing, thermophilic bacterium (strain KUI-RBT) was isolated from a geothermal spring biofilm in Rotorua, New Zealand. Strain KUI-RBT is a motile, straight rod, measuring approximately 0.7 {micro}m by 1.0 to 1.5 {micro}m with a diderm cell wall. Growth of KUI-RBT occurred from 39 to 74 {degrees}C (Topt 64.5 {degrees}C), pH 5.0 to 7.5 (pHopt 6.5), and 0 to 1% (w/v) NaCl (NaClopt 0.4-0.7%, w/v). KUI-RBT utilizes carbon dioxide and various organic carbon substrates as carbon sources and hydrogen as an electron donor. KUI-RBT can use oxygen (0-21%, v/v), elemental sulfur, thiosulfate, sulfite, nitrate, arsenate, and selenate as terminal electron acceptors. Major fatty acids of strain KUI-RBT include C20:1, C18:1, and C18:0 and the primary quinone is MTK-7. The whole genome G+C content is 34.23 mol%. Phylogenetic analyses indicate KUI-RBT to be a member of the family Hydrogenothermaceae, with Sulfurihydrogenibium azorense Az-Fu1T its closest characterised relative (94.51% 16S rRNA gene sequence similarity, 78.01% whole genome ANI, 61.34% whole genome AAI). Based on phylogenetic and phenotypic analyses, we propose KUI-RBT represents a novel genus and species within the family Hydrogenothermaceae, for which we propose the name Reysenbachia aerophila gen. nov., sp. nov. The type strain is KUI-RBT (=KCTC accession =JCM accession). The GenBank accession number for the 16S rRNA gene sequence of strain KUI-RBT is PZ052650. The GenBank accession number for the whole genome of strain KUI-RBT is JBVODP000000000.

3
Isolation and Genomic Characterization of Myxococcus faecalis Strains from Mangroves in Southeastern Brazil

Oliveira, R. S.; Lin, Y. F.; Jimenez, P. C.

2026-04-30 bioinformatics 10.64898/2026.04.28.721309 medRxiv
Top 0.1%
10.1%
Show abstract

Myxococcus faecalis was recently described from human fecal isolates, although subsequent evidence indicates an environmental distribution for this lineage. Here, we report the isolation and genomic characterization of two M. faecalis strains (BRX-014 and BRX-032) recovered from mangrove ecosystems along the southeastern coast of Brazil, representing the first record of the species in a marine-coastal biome. Phylogenomic reconstruction based on 120 conserved bacterial marker genes, together with Average Nucleotide Identity (ANI >97.6%) and digital DNA-DNA hybridization (dDDH 77.7-90.4%) analyses, confirmed their assignment to M. faecalis and demonstrated high genomic relatedness to strains previously recovered from soil and human feces samples. Pangenome analysis of five available genomes revealed a total repertoire of 9,827 genes, with a large core genome comprising 7,499 genes (76.3%), consistent with a highly conserved and nearly closed pangenome structure. Functional classification based on COG categories showed uniform distributions across all isolates. Comparative analysis of the degradome further revealed strong conservation of proteolytic and carbohydrate-active enzyme repertoires, dominated by serine and metallopeptidases and diverse glycoside hydrolases. The extensive genomic and functional similarity among isolates from geographically distant and ecologically distinct environments supports a broad ecological distribution of M. faecalis and suggests that its large and conserved genomic repertoire underpins its persistence across contrasting habitats. These findings expand the known ecological range of the species and provide a comparative genomic framework for future investigations into its distribution and functional potential across different habitats.

4
ERGA-BGE reference genomes of Hyalomma lusitanicum and its obligate Francisella endosymbiont as a genomic resource for One Health research

Uribe, J. E.; Echeverry-Perez, J. S.; Valcarcel, F.; Olmeda, A. S.; Sanchez-Sanchez, M.; Tercero, J. M.; Escudero, N.; Fernandez, R.; Boehne, A.; Monteiro, R.; Gut, M.; Aguilera, L.; Camara Ferreira, F.; Cruz, F.; Gomez-Garrido, J.; Alioto, T.; de Guttry, C.

2026-05-31 genomics 10.64898/2026.05.27.728183 medRxiv
Top 0.1%
9.8%
Show abstract

Hyalomma lusitanicum is a characteristic tick species of the western Mediterranean region, with a well-established distribution across the Iberian Peninsula. It is strongly associated with wild ungulates, particularly red deer, as well as livestock, to which it can transmit a wide range of pathogens, including viruses, bacteria, and protozoa. Here, we present three genomic resources for H. lusitanicum: a scaffold-scale nuclear genome, the complete mitochondrial genome, and the complete genome of its associated Francisella bacterial endosymbiont. The nuclear genome assembly spans 1.81 Gb and comprises 59 scaffolds, with a scaffold N50 of 153.6 Mb (L50 = 5) and no gaps, indicating high contiguity and completeness with a gene annotation completeness BUSCO score of 97.1 %. Genome annotation of the nuclear assembly identified 20,638 protein-coding genes, 1,422 non-coding genes, and 5,775 pseudogenes. A total of 18 scaffolds were assembled as putative chromosomes, exceeding the 11 chromosomes inferred as ancestral; however, synteny analyses suggest that several scaffolds likely represent fragmented portions of the same chromosome, probably due to incomplete Hi-C scaffolding. Despite this, the assembly represents one of the most complete tick nuclear genomes generated to date. In addition, we report the complete genome of a Francisella endosymbiont (1.51 Mb, 1,679 genes), characterized by a high proportion of pseudogenes and reduced genome size, consistent with patterns of genome reduction associated with obligate symbiosis. Together, these genomic resources provide a framework to investigate local adaptation and host-symbiont evolution, and to support improved surveillance, control, and management strategies for species of public health relevance.

5
Chromosome organization of Entamoeba histolytica and Entamoeba dispar

Kawano-Sugaya, T.; Kobayashi, S.; Kawashima, A.; Saito-Nakano, Y.; Izumiyama, S.; Nozaki, T.; Nakada-Tsukui, K.

2026-07-09 genomics 10.64898/2026.07.06.736064 medRxiv
Top 0.1%
8.8%
Show abstract

Entamoeba histolytica is a clinically important pathogenic eukaryote and the causative agent of amoebic dysentery. Entamoeba dispar, a nonpathogenic commensal species that resides in the human colon, is the closest sibling species, and serves as an appropriate comparator for genome-wide analysis. Although the genome of E. histolytica is approximately 26.9 Mb, and the largest known genome within the genus, that of E. invadens, is approximately 40.9 Mb, obtaining high-quality assemblies in this genus has remained challenging due to extensive repetitive regions, tRNA gene arrays, and aneuploidy. Here, we used PacBio HiFi sequencing to assemble the genomes of the pathogenic E. histolytica and the nonpathogenic E. dispar. We reconstructed all 36 chromosomes of E. histolytica and 35 chromosomes of E. dispar, assembling each as a single continuous DNA sequence (contig). The two species exhibited high genome-wide nucleotide similarity and conserved synteny at the amino acid level. At one end of each chromosome, we identified tRNA arrays, whereas the opposite end lacked such arrays, resulting in an asymmetric chromosomal architecture. Analysis of unique-read depth revealed widespread aneuploidy in both species: E. histolytica is predominantly tetraploid, whereas E. dispar is diploid, a conclusion further supported by SNP allele-frequency distributions. These assemblies provide a robust foundation for comparative genomics in Entamoeba and offer detailed insights into chromosome-end structure and ploidy.

6
Description of Rickettsia senegalensis sp. nov.: a new Rickettsia species detected worldwide

Labarrere, C.; Houmenou, C. T.; Fournier, P.-E.; Fenollar, F.; Mediannikov, O.

2026-05-05 microbiology 10.64898/2026.05.02.721834 medRxiv
Top 0.1%
8.0%
Show abstract

Rickettsia senegalensis is a novel Rickettsia species isolated from cat fleas, Ctenocephalides felis, in Senegal. Genomic analysis confirmed its status as a distinct species, placing it within the transitional Rickettsia group, within a R. felis cluster. Furthermore, rickettsial genes identical to those of Rickettsia senegalensis had been already identified in several hematophagous arthropods, including fleas and ticks parasitizing various hosts such as cats, dogs, opossums, and rodents in tropical and subtropical regions all over the world. It has also been detected in cat tissues, suggesting a potential host-pathogen association. Here we formally propose Rickettsia senegalensis sp. nov. as a new species. The type strain of this species is strain PU01-02T (= CSUR R184T = DSM 28250T). Strain PU01-02T grows aerobically in XTC-2, SF9, and LD652 cell lines at 28 {degrees}C in a CO2-free atmosphere. The genome of strain PU01-02T has a size of 1.62 Mb and a G+C content of 33.2%. RepositoriesThe genome sequence of Rickettsia senegalensis sp. nov. strain PU01-02T has been deposited in GenBank under accession number JBVYTQ000000000, and the rrs, gltA, ompB and sca4 gene sequences under accession numbers KF666476, KF666472, KF666470, KF666474, respectively. The plasmid accession numbers are PZ272915, PZ272916, and PZ272917, for pRS01, pRS02 and pRS03, respectively.

7
A gapless telomere-to-telomere reference genome of Ostreococcus tauri RCC4221 with expanded annotation of medium-sized ncRNAs

Liu, G.; Bousquet, L.; Mayeur, H.; Manirakiza, E.; Daric, V.; Klopp, C.; Noirot, C.; Lopez-Escardo, D.; Grimsley, N. H.; Yau, S.; Krasovec, M.; Echeverria, M.; PIGANEAU, G.

2026-07-14 genomics 10.64898/2026.07.10.737489 medRxiv
Top 0.1%
7.9%
Show abstract

Marine photosynthetic microbes contribute substantially to global primary production, yet many algal lineages still lack reference genomes with the continuity and annotation quality required for fine-scale structural, regulatory and comparative analyses. Ostreococcus tauri, one of the smallest known free-living photosynthetic eukaryotes, has been a model marine picoeukaryote for over two decades. Despite successive improvements to its historical reference genome, previous assemblies retained hundreds of gaps and incomplete genes, hampering high-resolution genomic analyses. Here, we present O. tauri RCC4221 genome version 2026, a telomere-to-telomere assembly of all 20 chromosomes spanning 13.34 Mb with no gaps. This assembly combines PacBio long-read sequencing, Illumina short-read polishing, correction of unresolved regions guided by independent Nanopore-based assemblies. The updated reference supports a curated annotation comprising 7,683 protein-coding genes, 48 tRNA genes, 3 rRNA operons, 116 medium-sized noncoding RNAs, one signal recognition particle RNA and 138 small nucleolar RNAs. It also improves gene-model integrity and recovers candidate coding loci absent from the 2014 reference. Structural analyses resolved the organization of the two atypical low-GC chromosome 2 and 19 that contain duplicated regions that were collapsed or misrepresented in previous assemblies. Finally, bisulfite sequencing and PacBio SMRT sequencing revealed a dual DNA methylation landscape, with CG-context cytosine methylation concentrated in gene bodies and N6-methyladenosine (m6A) enriched at the start codon. The updated O. tauri 2026 assembly provides a complete and curated reference resource for chromosome biology, comparative genomics, epigenomics and RNA biology in a model marine picoeukaryote.

8
Carbon monoxide utilisation by Thermanaeromonas species and description of Thermobium azorense gen. nov., sp. nov.

Galani, A.; Antony Venancius, M.; Tumulero, B.; Sipkema, D.; Sousa, D. Z.

2026-07-10 microbiology 10.64898/2026.07.10.736077 medRxiv
Top 0.1%
7.7%
Show abstract

Syngas fermentation by carbon monoxide (CO)-utilising acetogens offers a sustainable route for converting gasified waste materials into value-added chemicals. In this study, we isolated a novel thermophilic CO-utilising bacterium, strain AZ2, from marine hydrothermal sediment collected on the island of Sao Miguel, Azores, Portugal. Strain AZ2 is an obligately anaerobic, spore-forming bacterium. Average nucleotide identity (ANI; 78.4-86.7%) and digital DNA-DNA hybridization (dDDH; 23.4-32.5 %) analyses indicate that strain AZ2 represents a novel species within a previously uncharacterised lineage represented by the GTDB placeholder genus UBA2545 in the Neomoorellaceae family. Strain AZ2 was able to grow fermentatively on CO, producing acetate. We further demonstrated that its closest isolated relatives - Thermanaeromonas toyohensis, T. burensis, and Thermanaeromonas sp. strain 9S - are capable of growing on CO, producing either acetate or hydrogen gas (H2). Additionally, we unveiled the genomic potential for CO utilisation within other members of the GTDB placeholder class DSM-521 (previously Moorellia) to which our isolate belongs, expanding the list of possible thermophilic CO-utilising acetogens. We propose that strain AZ2T represents the type strain of a novel genus and species, named Thermobium azorense gen. nov., sp. nov. (= DSM 121889T = JCM 39698T).

9
The first reference genome assembly of the Chilean sea fig (Carpobrotus chilensis)

Lee, H.; D'Antonio, C. M.; Yi, S. V.

2026-06-19 genomics 10.64898/2026.06.15.732467 medRxiv
Top 0.1%
6.8%
Show abstract

Carpobrotus chilensis (Chilean sea fig) is a coastal succulent of uncertain origin that has naturalized along the California coast, where it co-occurs and hybridizes with the invasive species, Carpobrotus edulis. Despite their ecological importance and widely supported hybridization, genomic resources for this genus remain scarce. Here, we present a draft genome assembly of C. chilensis generated from PacBio HiFi long reads. The assembled nuclear genome spans 981.7 Mb across 178 contigs. The contig N50 was 73.0 Mb, and BUSCO completeness was 96.3%. K-mer and SNP-based analyses indicate extremely low heterozygosity (3.4 x 10-), reduced genetic diversity in this population. The genome is highly repetitive, with 81.67% of the sequences composed of transposable elements, predominantly long terminal repeat (LTR) retrotransposons. Gene prediction identified 21,744 protein-coding genes, with BUSCO completeness of 95.8%. Comparative analysis with C. edulis identified 8,783 single-copy orthologous gene pairs, with a median synonymous substitution rate (dS) of 0.019, indicating low sequence divergence between the two species. This genome assembly provides a foundational resource for investigating the genomic basis of hybridization and invasion in Carpobrotus.

10
Dengue virus in Solomon Islands 2023-2025: a whole genome surveillance study

Moselen, J.; Steinig, E.; Darcy, A.; Dofai, A.; Manele, A.; Lauri, B.; Joshua, C.; Mauruwai, P.; Aziz, A.; Horwood, P. F.; Orlando, N.; Caly, L.; Karan, N.; Solomon, J.; Lim, C. K.

2026-07-09 public and global health 10.64898/2026.06.30.26356797 medRxiv
Top 0.1%
6.8%
Show abstract

Background Following the cessation of COVID-19 travel restrictions in July 2022, concerns about a delayed dengue outbreak prompted the Solomon Islands Ministry of Health to establish enhanced genomic surveillance of circulating dengue virus (DENV) strains. Methods We performed amplicon-based whole genome sequencing (WGS) on PCR-positive serum samples collected at the National Referral Hospital, Honiara, between January 2023 and March 2025 (n = 63). Genomes were compared with publicly available sequences, and maximum-likelihood phylogenies were used to explore regional transmission dynamics. Findings We generated the first whole genome sequences from Solomon Islands (n = 45), with high recovery rates from acute infections (90%, Ct 17-44, mean coverage > 80%). Co-circulation of DENV-1, DENV-2, and DENV-4 was observed, with evidence of a serotype shift emerging in 2024. Phylogeographic analyses suggest ancestral introductions from Papua New Guinea for DENV-1 and DENV-2. Interpretation This study demonstrates the feasibility of whole genome sequencing for dengue surveillance in the Solomon Islands through a referral sequencing model that provides a pathway for progressive local capacity-building. By addressing technical challenges and critical gaps in regional genomic representation, our findings strengthen the evidence base needed for equitable and sustainable implementation of pathogen genomics across Pacific Island countries and territories.

11
Ruminococcus hollandia sp. nov. and Ruminococcus vasco sp. nov., two novel starch-degrading Ruminococcus isolated from the rumen of Holstein dairy cattle

Calapa, K. A.; Bock, R.; Embree, J.; LoBrutto, J.; Embree, M.

2026-07-11 microbiology 10.64898/2026.07.10.737842 medRxiv
Top 0.1%
6.3%
Show abstract

This study investigated the genomic and biochemical characteristics of two amylolytic microbial strains, NATIVEDY160T (= JE7B6T = NRRL B-68523T) and NATIVEDY161T (= JL13D9T, = NRRL B-68524T) isolated from the rumen of healthy Holstein dairy cattle. Both strains are obligately anaerobic, non-motile, Gram positive, catalase-negative, and oxidase-negative. Morphologically, NATIVEDY160T grows in long coccoid chains while NATIVEDY161T grows in short chains or pairs. NATIVEDY160T can catabolize amygdalin, esculin/ferric citrate, and starch, compared to NATIVEDY161T which utilizes amygdalin, arbutin, esculin/ferric citrate, glycogen, and D-maltose as determined by API 50 CH carbon panels. Starch degradation ability was verified for both strains, but neither showed cellulolytic activity as confirmed by starch agar and Congo red agar assays, respectively. HPLC analysis revealed that lactate was the primary end product of both strains carbohydrate fermentation, while strain NATIVEDY161T also produced small amounts of acetate. 16S rRNA sequences from both strains cluster with the Oscillospiraceae (formerly Ruminococcaceae) lineage Ruminococcus species, but average nucleotide identity of either strain compared to closely related Ruminococcus members was under the species threshold (95%). Genomic, phylogenetic, and phenotypic interrogation support NATIVEDY160T and NATIVEDY161T as novel species. Each strain was isolated from the rumen of dairy cows located within the central valley of southern California, which has a rich history of Dutch and Basque dairy farm ownership and is still the case today in the region. In recognition of the contributions and heritage of the central and southern California dairy industry, the names Ruminococcus hollandia and Ruminococcus vasco are proposed with NATIVEDY160T and NATIVEDY161T as their respective type strains.

12
Oligella otitidis sp. nov., isolated from middle ear discharge of children with chronic suppurative otitis media

Beissbarth, J.; Atto, B.; Mandal, P. K.; Cleanthous, A.; Harrison, B.; Gill, N. J.; Smith-Vaughan, H. C.; Kleinecke, M.; Rigas, V.; Leach, A. J.; Morris, P. S.; Marsh, R. L.

2026-06-30 microbiology 10.64898/2026.06.29.735399 medRxiv
Top 0.1%
5.6%
Show abstract

Oligella otitidis MSHR-50489EDL strain (ATCC: TSD462; DSMZ: DSM118617) is a new species of the genus Oligella that was isolated from a middle ear discharge swab from a child with chronic suppurative otitis media (CSOM). This Gram-negative coccobacillus produces small, circular, smooth, whitish-opaque and occasionally mucoid colonies. It grows in aerobic conditions at a temperature range from 25-42oC. Phylogenetic analysis demonstrates a relationship to other species of the genera Oligella and average nucleotide identity and digital DNA/DNA hybridization values indicate a distinct species in comparison to other Oligella species. Thus far, the majority of isolates exhibit resistance to ciprofloxacin, the first line treatment for CSOM.

13
Chromosome-level genome assemblies of the red algae Porphyra dioica and Porphyra linearis

Morcillo, J.; D hondt, S.; Lipinska, A.; Bouckenooghe, S.; Noyen, L.; Van de Vloet, A.; Vranken, S.; Knoop, J.; Leliaert, F.; De Clerck, O.

2026-05-16 genomics 10.64898/2026.05.14.725108 medRxiv
Top 0.1%
5.5%
Show abstract

As one of the earliest-diverging multicellular eukaryotic lineages, the bladed Bangiales (Rhodophyta) possess a deep evolutionary history with a central role in the multi-billion-dollar global seaweed aquaculture industry. Although North Atlantic representatives are emerging candidates for regional mariculture, the scarcity of high-quality genomic resources for these taxa hinders both fundamental research and commercial optimization. To address this, we present the first chromosome-level genome assemblies for two native European species: Porphyra dioica (150.44 Mbp) and Porphyra linearis (95.22 Mbp). By integrating Oxford Nanopore Technologies (ONT) long-read sequencing with Hi-C proximity ligation, we generated highly contiguous nuclear genomes resolved into five chromosomes. Structural gene models were predicted through the BRAKER3 pipeline, identifying 12,548 and 10,382 protein-coding genes for P. dioica and P. linearis, respectively. Subsequent homology-based functional annotation characterized 57.4% and 59.8% of these predicted proteins. Supplemented by circularized organellar genomes, these reference genomes provide a critical framework for future research, enabling comparative studies of Atlantic-Pacific divergence and facilitating the development of selective breeding programs for sustainable European aquaculture.

14
Phylogenomic description of three novel species of the Microbulbifer genus, phylum Pseudomonadota, isolated from marine sponges and corals

Tang, Y.; Track, A.; Miller, N. A.; Mandelare-Ruiz, P.; Paul, V. J.; Konstantinidis, K. T.; Agarwal, V.

2026-06-11 microbiology 10.64898/2026.06.10.731415 medRxiv
Top 0.1%
5.4%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWUnderstudied bacterial genera present a dynamic phylogenetic landscape and opportunities for discovering new taxa as more strains are isolated and genomic data is added. Here, through phylogenomic analysis, we describe three novel species of the globally distributed cosmopolitan marine bacterial genus Microbulbifer. This genus is ubiquitous in saltwater microbiomes and is a validated source of biodegradation enzymes as well as high value small molecule natural products. Average nucleotide identity (ANI) to the closest known species, Microbulbifer variabilis ATCC 700307T, was less than 88.4% for all three novel species. Isolates of the three novel species, designated as PAAF003T (T = type strain), ZKSA006T, and SSSA003T were imaged to reveal their phormological characteristics. Based on phylogenetic data, strains PAAF003T, ZKSA006T, and SSSA003T represent three new species of the genus Microbulbifer, for which the names Microbulbifer maximicatervae sp. nov., Microbulbifer regidiadema sp. nov., and Microbulbifer mixtoriginis sp. nov. are proposed, respectively, under the SeqCode. We also reconstructed a robust phylogeny of available Microbulbifer genomes, which should faciliatate future isolation and strain description studies.

15
ArchaeaHQ: A Curated Reference Database of Archaeal Genomes

Bespiatykh, D.; Leao, P.

2026-06-03 microbiology 10.64898/2026.06.02.729493 medRxiv
Top 0.1%
5.4%
Show abstract

Archaea have proven to be major players in biogeochemical cycles across diverse ecosystems, yet we still see an underrepresentation of archaeal genomes in the datasets used by popular computational biology tools. Here we present ArchaeaHQ, a quality-controlled, systematically curated reference database of 21,644 archaeal genomes compiled initially from 35,993 assemblies from all four archaeal kingdoms retrieved from NCBI: Methanobacteriati (Euryarchaeota), Thermoproteati (TACK), Nanobdellati (DPANN), and Promethearchaeati (Asgard). All genomes in the database passed standardized quality control, requiring [≥]70% completeness and [≤]10% contamination. A total of 44.2% of genomes in ArchaeaHQ achieved [≥]90% completeness, while 93.1% exhibited [≤]5% contamination. ArchaeaHQ comprises 16,199 metagenome-assembled genomes (MAGs; 74.8%) and 5,445 isolate genomes (25.2%). Approximately 75% of MAGs are assigned to 17 ecologically meaningful categories based on sampling origin, and around 65% of genomes include geographic metadata. ArchaeaHQ is available at https://doi.org/10.6084/m9.figshare.32266599 and provides an analysis-ready reference set for metagenomic classification, biogeochemical and ecological studies, comparative genomics, and development of archaeal-specific bioinformatic tools. Impact StatementArchaea are key drivers of the global carbon, nitrogen and methane cycles, yet their genomes remain underrepresented and inconsistently curated in the public databases that power modern computational biology tools. We present ArchaeaHQ, a quality-controlled, systematically curated reference set of 21,644 archaeal genomes spanning all four archaeal kingdoms, each passing standardized completeness and contamination thresholds and enriched with environmental and geographic metadata. By providing an analysis-ready, downloadable resource compatible with standard pipelines, ArchaeaHQ fits the gap between taxonomy-focused frameworks and unfiltered genome archives supporting metagenomic classification, biogeochemical and ecological studies, comparative genomics, and the development of archaeal-specific bioinformatic tools.

16
Highly contiguous reference genome assembly of the endangered Orces blue whiptail Holcosus orcesi

Pozo, G.; Cisneros-Heredia, D. F.; Barragan-Orbe, D.; Sanchez-Nivicela, J. C.; Arbelaez, E.; Torres, M.

2026-05-16 genomics 10.64898/2026.05.14.725226 medRxiv
Top 0.1%
5.2%
Show abstract

Holcosus orcesi, the Orces Blue Whiptail, is a Critically Endangered lizard endemic to the upper Jubones River basin in southern Ecuador. Restricted to a narrow elevational range within semi-arid Andean shrublands, it represents one of the few montane members of a predominantly lowland lineage. Here we present the first high-quality reference genome for H. orcesi, generated using Oxford Nanopore Technologies long-read sequencing. The assembly spans 1.68 Gb across only 91 contigs, with an N50 of 76.2 Mb and a BUSCO completeness of 96.8%, making it among the most contiguous and complete squamate genomes to date. Structural annotation predicted 25,682 genes, of which 85% showed homology to known proteins and 45% were assigned Gene Ontology terms. Repetitive elements accounted for 46.3% of the genome, with LINEs representing the predominant class. This genome provides a foundational resource for future evolutionary, comparative and conservation-genomic research of H. orcesi and other mountain reptiles, enabling studies of population genomics, local adaptation, and genomic erosion in isolated populations. By expanding the genomic representation of tropical montane reptiles, this work helps address longstanding phylogenetic and geographic gaps in global biodiversity genomics and provides a foundation for evidence-based conservation of H. orcesi and related taxa.

17
Biotechnological potential of aromatic compounds utilizing bacteria from Brazilian caves, including a novel cave Nocardioides sp. SF1

Marques, E. d. L. S.; Gross, E.; Jambeiro, I. C. d. A.; Souza, M. C. B.; Dias, J. C. T.; Rezende, R. P.

2026-06-24 microbiology 10.64898/2026.06.23.734003 medRxiv
Top 0.1%
4.3%
Show abstract

From Brazilian limestone caves, we isolated 29 bacteria utilizing phenol (23 bacteria), toluene (all bacteria), and/or benzene (all bacteria) as sole carbon sources. One isolate showed phosphate solubilization, while lipase/esterase activity occurred in two isolates; no amylase activity was detected, but 16 isolates ([~]55%) exhibited protease activity. Among them, Nocardioides sp. SF1 was selected for whole-genome sequencing due to its aromatic compound tolerance and protease activity. Additionally, catechol cleavage assays yielded unexpected purple pigmentation, suggesting non-canonical aromatic metabolism. Its high-quality draft genome (4.25 Mbp, 16 contigs, N50 of 887 kb) lacks canonical phenol hydroxylase but encodes alternative oxidation systems, phenylacetyl-CoA pathway, besides, desferrioxamine siderophore, biosurfactants, and phosphate solubilization, key adaptations for oligotrophic caves and biotechnologically interesting activities. Whole-genome comparisons (TYGS/GGDC, OrthoANI and k-mer) suggest potential new species. Lacks acquired antimicrobial resistance genes (ResFinder) and pathogenicity potential (PathogenFinder). Nocardioides sp. SF1 emerges as a non-pathogenic candidate for aromatic bioremediation and plant growth promotion in contaminated, nutrient-poor environments, highlighting cave actinobacterias unexplored biotechnological potential.

18
First outbreak of Lumpy Skin disease in Catalonia, Spain, 2025-2026

Obregon-Gutierrez, P.; Correa-Fiz, F.; Fonseca-Rodriguez, O.; Cortey, M.; Cobos, A.; Riera, C.; Soler, M.; Ribas, N.; Domenes, F.; Pailler-Garcia, L.; Domingo, M.; Majo, N.; Vidal, E.; Lorca-Oro, C.

2026-06-22 genomics 10.64898/2026.06.18.733166 medRxiv
Top 0.1%
4.0%
Show abstract

Lumpy skin disease (LSD) is an emerging cattle disease caused by lumpy skin disease virus (LSDV), with major impacts on the industry, being classified as a Category A disease. Although it was historically confined to Africa, LSD has expanded into the Middle East, Asia and Europe. Here, we report two LSDV genomes from the first outbreak detected in Catalonia, Spain, in October 2025. The genomes were assembled from high-throughput sequencing data generated from two homogenized skin nodules. Comparative phylogenetic analyses were performed using all available complete LSDV genomes and rpo30 gene sequences. These analyses placed the LSDV isolates detected in Catalonia within clade 1.2, closely related to the isolates recently reported in Sardinia, Italy. Our findings also support a connection between recent south-western Europe and central African strains, possibly through northern Africa, and highlight the need for more complete genomes to clarify the origin and connections among recent LSDV outbreaks.

19
Phylogenomic Taxonomic Analysis of Ralstonia solanacearum Strains causing Bacterial Wilt Disease in Northeastern Argentina.

Obregon, V.; Shin, G. Y.; Galdeano, E.; Escobar, R.; Lattar, T.; Ibanez, J. M.; Amadio, A.; Irazoqui, J. M.; Santiago, G. M.; Eberhardt, M. F.; Gochez, A. M.; Lowe-Power, T.

2026-05-01 microbiology 10.64898/2026.04.29.721750 medRxiv
Top 0.1%
4.0%
Show abstract

Ralstonia solanacearum species complex (RSSC) is a genetically diverse group of plant pathogens, yet genomic data from South America remain limited. Here, we characterize 13 RSSC strains isolated from tomato, pepper, and eggplant in northeastern Argentina. Phylogenetic analysis of the egl marker gene assigned these strains to phylotype IIA and suggested two closely related lineages. Complete genomes (5.63-5.76 Mb) were generated for four representative strains, yielding high-quality (99.94% completeness with f_Burkholderiaceae CheckM markers), closed assemblies with canonical bipartite architecture. Phylogenetic analysis of the egl marker, 49 conserved bacterial genes, and average nucleotide identity (ANI) analyses, consistently assigned one lineage to sequevar IIA-50, forming a coherent and monophyletic group. In contrast, although egl analysis suggested the second lineage was related to one sequevar IIA-38 reference strain, genomic analysis did not support this assignment. Further, the genomic analysis revealed significant genomic distance between the genomes for two sequevar 38 representative strains, supporting a conclusion that sequevar 38 itself was not monophyletic and instead appears paraphyletic. These findings highlight limitations of single-locus classification and support genome-informed refinement of RSSC sub-phylotype taxonomy. Outcome statementReports of bacterial wilt disease in Argentina had not yet been published in the international literature although the disease has been long-standing. This study provides complete genome sequences for four Ralstonia solanacearum strains from Northern Argentina and places them within a global phylogenomic framework. The Argentine strains cluster into two closely related phylotype IIA lineages, indicating that bacterial wilt in this regional dataset is associated with genetically similar populations. For clear communication of which strains are present in Northern Argentina, we attempted to classify the lineages to the long-standing sequence variant (sequevar) system for naming R. solanacearum species complex (RSSC) strains. One lineage was confidently assigned to IIA-50 with genomic support that confirmed phylogenetic analysis of the classical genetic marker egl. However, newly available genomes for sequevar reference strains revealed an issue where two distantly related strains are currently recognized as references for sequevars. Overall, these results provide evidence supporting the need for genome-informed refinement of sub-phylotype classification and expand genomic representation of South American RSSC populations. Data summaryComplete genome assemblies and raw reads for INTABV18, INTABV29, INTABV624 and INTABV2657 are deposited to NCBI under the project number PRJNA1407867. The curated dataset of public RSSC genomes is available to users who register a free account on KBase via a KBase narrative (https://narrative.kbase.us/narrative/189849). The narrative described in a living BioRxiv pre-print [1]. Supplemental files such as Figure S1, rectangular versions of all trees (Figure 2 and 3 and S1) and supplementary table S1, S2, S3 and S4 are available on Zenodo at doi.org/10.5281/zenodo.19502890 O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=172 SRC="FIGDIR/small/721750v1_figS1.gif" ALT="Figure 1"> View larger version (47K): org.highwire.dtl.DTLVardef@1ac3168org.highwire.dtl.DTLVardef@1dfd0d6org.highwire.dtl.DTLVardef@107ae42org.highwire.dtl.DTLVardef@141937c_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure S1.C_FLOATNO Maximum-likelihood phylogenetic tree inferred from 471 bp of endoglucanase (egl) gene sequences assigned Argentine strains as phylotype II sequevar 38 and sequevar 50. The tree was constructed using PhyML v3.0 under the GTR nucleotide substitution model with gamma-distributed rate heterogeneity ( = 0.33), as selected by the SMART model selection procedure implemented in PhyML (Lefort et al., 2017). The egl sequences from Argentine strains are highlighted in blue, and their corresponding GenBank accession numbers for both the egl nucleotide sequence and the whole-genome assembly are shown in parentheses. Reference egl sequences representing sequevars IIA-38 (CFBP6801 and CIP120) and IIA-50 (T1-UY and ACH1076) are also shown in bold and marked with yellow circles. A searchable PDF of this tree in rectangular format is available on Zenodo (doi.org/10.5281/zenodo.19502890). C_FIG O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=196 SRC="FIGDIR/small/721750v1_fig2.gif" ALT="Figure 2"> View larger version (53K): org.highwire.dtl.DTLVardef@39d776org.highwire.dtl.DTLVardef@170bd89org.highwire.dtl.DTLVardef@aba166org.highwire.dtl.DTLVardef@1f156dd_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 2.C_FLOATNO Maximum-likelihood phylogenetic tree inferred from 710 bp of endoglucanase (egl) gene sequences assigned Argentine strains as phylotype II sequevar 38 and sequevar 50. The phylogenetic tree was constructed using PhyML v3.0 under the GTR+R nucleotide substitution model, as selected by the SMART model selection procedure (Lefort et al., 2017). egl sequences from four Argentine strains (INTABV18, INTABV29, INTABV624, and INTABV2657) are shown in bold and highlighted in blue. Reference egl sequences representing sequevars IIA-38 (CFBP6801 and CIP120) and IIA-50 (T1-UY and ACH1076) are also shown in bold and marked with yellow circles. Two USA strains identified as IIA-38 (UCD576 and RS124) are shown in bold. A searchable PDF of this tree in rectangular format is available on Zenodo (doi.org/10.5281/zenodo.19502890). C_FIG O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=116 SRC="FIGDIR/small/721750v1_fig3.gif" ALT="Figure 3"> View larger version (37K): org.highwire.dtl.DTLVardef@17dd372org.highwire.dtl.DTLVardef@1c5156corg.highwire.dtl.DTLVardef@179d9org.highwire.dtl.DTLVardef@e6d529_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 3.C_FLOATNO Approximate maximum-likelihood phylogeny based on a concatenated alignment of 49 conserved genes places four Argentine genomes (INTABV18, INTABV29, INTABV624 and INTABV2657) within the phylotype IIA clade. The tree was constructed using the SpeciesTreeBuilder v0.1.4 application on the KBase platform, incorporating the four Argentine genomes into a reference dataset of 825 genomes representing the known global diversity of the RSSC. The tree was visualized and annotated using iTOL v7.4.2. Argentine genomes are shown in bold and highlighted in blue, and egl reference strains for the sequevar IIA-38 (CIP120 and CFBP6801) and IIA-50 (T1-UY) are shown in bold and marked with yellow circles. Branches with approximate likelihood-ratio support values higher than >70% are colored in blue. A searchable PDF of this tree in rectangular format is available on Zenodo (doi.org/10.5281/zenodo.19502890). C_FIG

20
Genomic Decoding of Specialized Aromatic Hydrocarbon Degradation in Mangrove-Derived Gordonia sp. B7 2

Jiang, F.; Shi, H.; Lu, M.; Zhao, Z.; Xu, X.; Feng, H.

2026-05-25 microbiology 10.64898/2026.05.23.727409 medRxiv
Top 0.1%
3.4%
Show abstract

Petroleum pollution has increased worldwide, driving the search for microorganisms with efficient hydrocarbon-degrading capabilities. Here, we report a novel bacterium, Gordonia sp. B7-2, isolated from mangrove sediments in Hainan, China. Phylogenetic analysis based on the 16S rRNA gene and whole-genome sequences, together with digital DNA-DNA hybridization and average nucleotide identity values, supported its classification as a new species within the genus Gordonia. The complete genome of strain B7-2 consists of a single circular chromosome of 5.39 Mb with a G+C content of 65.99%, and encodes 4,887 protein-coding genes. Genomic annotation revealed a complete pathway for aromatic hydrocarbon degradation, including genes encoding protocatechuate 3,4-dioxygenase and biphenyl-2,3-diol 1,2-dioxygenase, whereas genes involved in the initial oxidation of alkanes were absent. Consistent with these genomic predictions, strain B7-2 degraded 64.33% of crude oil (300 mg/L) within 28 days, with rapid degradation during the initial 14 days, followed by a slower phase thereafter, reflecting the dynamics of complex hydrocarbon mixtures. Together, these results demonstrate that strain B7-2 is specialized for the degradation of aromatic hydrocarbons and highlight its potential for targeted petroleum bioremediation. IMPORTANCEMangrove ecosystems are highly vulnerable to petroleum contamination, yet the microorganisms responsible for hydrocarbon turnover in these environments remain poorly characterized. This study describes a new bacterial species, Gordonia sp. B7-2, that exhibits a strong metabolic specialization for aromatic hydrocarbon degradation. Unlike many known oil-degrading bacteria that preferentially utilize alkanes, strain B7-2 targets aromatic components of crude oil, which are among the most persistent and toxic fractions.Its ability to efficiently degrade crude oil highlights its potential in the bioremediation of contaminated coastal environments and expands our understanding of microbial contributions to hydrocarbon cycling in mangrove sediments